Medical Image Analysis
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Medical Image Analysis's content profile, based on 35 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Chen, J.; Pham, T.-H.; Zhang, P.; Varghese, J.
Show abstract
Accurate measurement of intra-cardiac blood oxygen (O2) saturation is essential for cardiovascular assessment, yet current methods require invasive catheterization. T2-based cardiac magnetic resonance imaging (CMRI) enables non-invasive O2 quantification, but deep learning automation is constrained by scarce annotated data. We propose a unified self-supervised learning (SSL) framework integrating cine CMRI and T2 oximetry CMRI to learn generalizable representations without labels. Our approach pre-trains ResNet and vision transformer encoders using contrastive learning and masked image modeling on over 48,000 cardiac images. Pre-trained encoders are fine-tuned for O2 saturation regression with uncertainty quantification to enhance clinical trustworthiness. Our SSL framework significantly outperforms traditional radiomics and supervised baselines, with SimCLR pre-trained ResNet achieving a mean absolute error of 3.70, representing over 15\% improvement. These findings demonstrate SSL's potential to address annotation bottlenecks in non-invasive cardiac diagnostics.
Kamalakannan, N. K.; Kamalakannan, J.
Show abstract
Deep segmentation networks can degrade sharply when an expected MRI sequence is unavailable at inference. We present NeuroMesh, a bottleneck controller that combines a gated recurrent unit (GRU) with a graphconvolutional edge-activation mask, designed to adapt a U-Net-style segmentation backbone to missing input. We evaluate NeuroMesh in a pilot study using a 30-patient subset of the BraTS 2020 benchmark (22 training, 4 validation, and 4 held-out test patients) under a prespecified frozentest protocol. On the frozen test set, NeuroMesh has higher tumor-core and enhancing-tumor Dice than a plain U-Net in most evaluated missing-modality conditions, but wholetumor Dice falls from 0.596 to 0.108 when FLAIR is missing, compared with 0.604 to 0.545 for the plain U-Net. Direct analysis of the predicted edge-activation mask shows negligible change across modality-availability conditions. A parameter-light static-gating control reproduces the FLAIR failure mode without recurrence, a failure-signal input, or graph-structured machinery. These results do not support the intended interpretation that the trained controller performs input-conditional topology rewiring at the scale of this pilot. Instead, they expose a discrepancy between architectural intent and realized behavior and identify a specific missing-modality failure mode that warrants further investigation. Given the small validation and test sets, the findings are descriptive and do not establish clinical or population-level generalization.
Zheng, J.; Kalaie, S.; Ma, Q.; Meng, Q.; Rjoob, K.; Gifani, P.; Hu, L.; Babazade, N.; Coriano, M.; Zhong, W.; Vafaeezadeh, M.; Tahasildar, S.; Vadgama, N.; Senevirathne, D. S.; Santhirasekaram, A.; McGurk, K. A.; Curran, L.; He, Y.; Chen, L.; Mo, Y.; Huang, L.; Qiao, M.; Huang, Y.; Bai, W.; O'Regan, D. P.
Show abstract
Cardiac imaging enables quantitative assessment of cardiac structure and function but remains constrained by cost, infrastructure and specialist expertise. In contrast, electrocardiogram (ECG) is widely accessible yet underexploited, despite encoding latent information about cardiac physiology. Here we introduce visionECG, a conditional flow matching framework that learns a probabilistic mapping between two biological distributions - the space of cardiac electrical signals and the space of cardiac geometries. Using 71,132 paired ECG and cardiac mesh sequence datasets from the UK Biobank, with external assessment in 5,000 patients with ECG-echocardiogram pairs, the model reconstructs quantitatively accurate spatiotemporal representations of the left ventricle using ECG inputs and basic demographic information alone. These reconstructions enable discrimination of structural abnormalities and disease labels, provide visualisations of functional abnormalities, and support flexible quantification of both global and regional parameters. By reframing the ECG as a generative source of patient-specific left ventricular geometry and motion, this work establishes a scalable framework for translating low-dimensional signals into high-dimensional, physiologically grounded structured representations.
Mukhopadhyay, A.; Halder, K.; Neogy, R.
Show abstract
Mapping hierarchical brain networks within traditional Euclidean space causes significant structural distortion, undermining neuroimaging diagnostic frameworks. While hyperbolic models like the Poincare ball preserve these nested topologies, they demand heavy computational overhead due to intricate Mobius operations and curved geodesics. This paper introduces a highly efficient non-Euclidean framework for analyzing neurocognitive decline utilizing the Beltrami-Klein ball model. By projecting hyperbolic geodesics as Euclidean straight lines, this approach converts complex distance calculations into simple dot products, radically reducing processing demands. We validated our methodology against state-of-the-art Poincare and Lorentz baselines using datasets for Schizophrenia, Parkinsons Disease, and Alzheimers Disease. The Klein-based framework demonstrates superior performance, delivering both higher diagnostic precision and accelerated processing velocities across all three neurocognitive disorders.
Tahmasebidehkordi, H.; Bahramy, A.; Julian, D. R.; Cohen, J. A.; Neal, M.; Bumgardner, C.; Nelson, P. T.; Pearce, T. M.; Kofler, J.
Show abstract
IntroductionCerebral amyloid angiopathy (CAA) is characterized by amyloid-beta deposition in cortical and leptomeningeal vessels and associated with cognitive impairment and hemorrhage. Current neuropathological assessments rely on semiquantitative grading and lack vessel-level resolution and scalability. Existing computational pathology approaches also fail to capture individual vessel morphology and spatial amyloid distribution across whole-slide images (WSIs). To address this gap, we developed a deep learning framework for reproducible, quantitative analysis of CAA in WSIs. MethodsWe analyzed 20 postmortem brain tissue sections from the frontal (n = 10) and occipital cortices (n = 10) of 10 individuals with Alzheimers disease pathology obtained from the University of Pittsburgh Alzheimers Disease Research Center, which served as the internal development cohort. An independent external cohort consisted of 10 sections (5 frontal and 5 occipital samples) from 5 individuals obtained from the University of Kentucky Alzheimers Disease Research Center. We trained and compared three semantic segmentation architectures, a standard U-Net, a dual-attention residual U-Net (DA-ResUNet), and a Swin Transformer-based U-Net (Swin-UNet), using the internal development cohort with slide-level five-fold cross-validation. All models were evaluated on the independent external cohort to assess generalization under domain shift. Based on segmentation performance and computational efficiency, we selected one architecture to generate whole-slide composite segmentation masks for vessel walls, amyloid deposits, and tissue compartments. These masks were subsequently used for deterministic vessel detection, morphometric measurements, and quantification of vascular and perivascular amyloid features through post-processing analysis. ResultsAll three architectures achieved high segmentation accuracy on the internal cohort, with Dice scores above 90% across vessel walls, amyloid deposits, gray matter, and leptomeninges. The Swin-UNet showed marginally higher performance for vessel segmentation, whereas the DA-ResUNet provided more balanced accuracy and computational efficiency and was selected for downstream analysis. External cohort evaluation demonstrated robust generalization, with attention-enhanced models outperforming the standard U-Net under domain shift. Using the selected model, the pipeline reliably detected valid vessels, excluded non-vascular artifacts, and enabled deterministic extraction of vessel morphometry, vascular and perivascular amyloid burden, and identification of circumferential CAA involvement at the vessel level. DiscussionThis framework provides a scalable, interpretable solution for vessel-level CAA analysis, supporting robust geometric and spatial characterization of cerebrovascular pathology and enabling future integration with clinical and genetic studies. Beyond CAA, the modular design allows extension to other vascular pathologies, including arteriolosclerosis, in WSIs, facilitating broader investigation of cerebrovascular disease mechanisms.
Bit, S.; Guney, O. B.; Jia, S.; Kolachalama, V. B.
Show abstract
Automated interpretation of neuroimaging studies requires simultaneous assessment of multiple imaging evidence variables, each tied to distinct anatomical structures. Vision-language models (VLMs) offer a unified framework for multi-task analysis, but adapting pre-trained VLMs remains challenging. Full fine-tuning is computationally prohibitive, and joint multi-task training requires simultaneous access to all task data, which is often infeasible in clinical settings. Although model merging enables multi-task composition without joint re-training, existing methods focus on post-hoc algorithms with limited extension to VLMs and minimal application to neuroimaging. Here, we present GRadient-guided Adapter Merging (GRAM), a layer-selective low-rank adaptation (LoRA)-based fine-tuning and merging framework for multi-task neuroimaging visual question-answering (VQA). GRAM uses a gradient ratio that contrasts class-specific gradients to identify task-discriminative layers, and applies subspace-constrained projected gradient descent to restrict LoRA updates to directions consistent with the geometry of the pre-trained model. We leveraged a structured VQA benchmark, developed from the National Alzheimer's Coordinating Center (NACC) dataset, that pairs multi-sequence brain MRI studies with question-answer pairs across clinically relevant imaging evidence variables. Experiments on the VQA benchmark showed that GRAM outperformed or matched all-layer LoRA fine-tuning and a standard merging baseline while reducing inter-task interference during merging, and approached or surpassed the performance of joint multi-task training without joint re-training.
Li, Z.; Sun, Y.; Jiang, C.; Pan, T.; Zhou, Y.; Wang, C.; Pan, L.; Zhang, X.; Yang, Z.; Yu, Z.; Xiao, Z.; Chen, J.; Huang, Y.; Sun, R.; Gan, Y.; Li, X.; Zhang, B.; Zhang, Z.; Wang, X.; Han, L.; Qi, Y.; Cheng, Y.; Liang, Y.; Ge, J.
Show abstract
BACKGROUND: Coronary angiography remains the reference standard for diagnosing coronary artery disease and guiding revascularization, yet its interpretation requires expert integration of multi-view anatomy, lesion morphology and procedural context. Existing artificial intelligence approaches are largely task-specific, annotation-dependent and limited in capturing the semantic relationship between angiographic findings and interventional decision-making. Whether large-scale vision-language pretraining can enable transferable foundation-model representations for invasive coronary imaging remains unknown. METHODS We developed CAG-MIND, a domain-specific vision-language foundation model for coronary angiography, using 135,475 CAG examinations paired with procedural reports, comprising 812,850 angiographic videos from Zhongshan Hospital and Shanghai Geriatric Medical Center. Each case consisted of standardized six-view angiographic acquisitions paired with structured procedural semantics extracted from routine reports using a large language model-assisted pipeline. The model was pretrained by aligning multi-view angiographic representations with report-derived semantic embeddings through bidirectional contrastive learning. Performance was evaluated under zero-shot and supervised fine-tuning settings across 11 downstream tasks grouped into structural abnormality detection, atherosclerotic plaque assessment, and interventional decision prediction, using both an internal validation cohort and an independent external test cohort. RESULTS CAG-MIND demonstrated robust performance across all three task categories. In the zero-shot setting, the model achieved mean AUROCs of 0.686 in the internal validation cohort and 0.745 in the external test cohort, indicating transferable multimodal representations without task-specific supervision. Following supervised fine-tuning, the mean AUROC increased to 0.827 and 0.846, respectively, with excellent performance for coronary stenosis detection (AUROC 0.940 in both cohorts), balloon/stent prediction (0.900 and 0.907), and CABG recommendation (0.877 and 0.875). Compared with representative biomedical vision-language models and conventional image-based architectures, CAG-MIND consistently achieved superior performance in both zero-shot and supervised settings and remained superior to fully fine-tuned competing models when trained with only 10% of the labelled data. Grad-CAM visualization demonstrated anatomically plausible lesion-focused attention, supporting the interpretability of the learned representations. CONCLUSIONS CAG-MIND is, to our knowledge, the first large-scale vision-language foundation model for coronary angiography trained at more than 100,000-patient scale. By aligning standardized multi-view angiographic videos with report-derived procedural semantics, CAG-MIND enables robust zero-shot transfer, data-efficient fine-tuning and cross-center generalization. These findings support domain-aligned multimodal pretraining as a scalable foundation-model paradigm for invasive cardiovascular imaging and future cath-lab decision support.
Chattopadhyay, T.; Shelar, K.; Thomopoulos, S. I.; Thompson, P. M.
Show abstract
Scaling laws describe how model performance improves as the amount of training data increases, and recent theories such as the zeta law suggest that scaling behavior is influenced by the eigenspectrum of the models latent representation. Here, we evaluated whether the distribution of discriminative signals across spectral modes predicts the future scaling behavior, for MRI transformers trained for disease classification. We trained three supervised 3D vision transformers (ViT3D, MINiT, and NIT) for Alzheimers disease classification using 2,822 training scans from the Alzheimers Disease Neuroimaging Initiative (ADNI); we compared their encoder spectra with that of a frozen self-supervised DINO ViT-B/16 encoder adapted to 3D MRI. The supervised models learned highly concentrated representations, with 90-96% of CLS-token variance captured by a single principal component, whereas DINO distributed signal across many latent directions. Via spectral expansion of the Mahalanobis signal, we found that supervised training concentrated disease information into a single dominant mode, while self-supervised training produced a richer spectral geometry with higher effective rank and discoverability. This led to different scaling behavior: supervised models exhibited flatter AUC(N) curves, yet DINO continued to improve as sample size increased, gaining 11.0 percentage points from N=50 to N=2,822. Overall, the spectral distribution of the discriminative signal, for these different encoder types, influenced how much performance remained discoverable as sample size increased. Distributed representations may retain signal across many latent modes and continue to improve with additional data, whereas concentrated representations tend to exhaust most of the discoverable signal at much lower sample sizes.
Kuruba, S.;Stephenson, G.;Kasinath, V.
Show abstract
Multi-organelle segmentation in volumetric electron microscopy (vEM) faces several challenges, including severe class imbalance, the presence of small, rare classes, and inconsistent class coverage across crops. While recent work has focused primarily on architectural design, the impact of sampling, loss functions, and masking strategies on training effectiveness remains comparatively underexplored in vEM organelle segmentation. Here, we systematically evaluate sampling strategies, loss configurations, masking approaches, and model families (CNNs and vision transformers) on the CellMap benchmark. Using 289 annotated 3D FIB-SEM crops, we establish a 32-class segmentation benchmark with stratified train, validation, and test splits, and evaluate all the methods under the same training and inference settings. Across controlled ablations, the proposed combination of repeat-factor sampling, Tversky-BCE loss, and masking achieved the strongest rare-class performance, increasing rare-class mean Dice (mDice) from 0.3244 under uniform sampling to 0.3409. This corresponds to an absolute gain of +0.0165 mDice and a 5.1% relative improvement, while preserving comparable performance on common classes. Overall, we find that sampling, loss design, and masking contribute as much to performance variation as the choice of architecture, highlighting the importance of training-recipe design alongside model architecture in vEM organelle segmentation.
Miri Rekavandi, A.; Jbabdi, S.; Smith, S. M.
Show abstract
This paper presents a framework for modelling the topography of whole-brain connectivity in resting-state functional MRI. The aim is to disentangle functional segregation, which manifests as abrupt changes in connectivity, from so-called gradients, i.e., smooth variations in connectivity across the brain. Our core assumption is that functional segregation leads to low-rank structure in the dense (point-to-point) connectome, whereas connectivity gradients imply a sparse and non-low-rank structure in the dense connectome. Our method thus decomposes the connectome into low-rank and sparse components, enabling the integration of local-nonlinear and global-linear embedding strategies. We show that this hybrid model approximates the empirical dense connectome more effectively than purely low-rank or purely gradient approaches. We also find that connectivity gradients derived from this model exhibit strong correspondence with task-based topographic maps. We hope that this approach can provide insight into the organisational principles of brain regions where gradients remain poorly characterised.
Hasny, M.; Daza, L.; Bressem, K.; Di Folco, M.; Schnabel, J. A.
Show abstract
Early and accurate risk stratification of cardiovascular disease (CVD) is crucial to initiate timely preventive interventions. As large-scale multimodal clinical cohorts become increasingly available, there is growing interest in whether incorporating additional sources of information can improve CVD risk stratification. Cine cardiac MR (CMR) represents a compelling example of such a source, as it captures objective, high-dimensional structural and functional information about the heart, independent of patient-reported data. In this study, we deploy a flexible vision-tabular method to incorporate cine CMR into CVD risk assessment together with structured clinical data. Using a large prospective imaging cohort from the UK Biobank, we show that cine CMR encodes CVD risk beyond established risk scores, increasing AUROC by 0.036 over SCORE2, the best-performing traditional risk score (0.742 vs. 0.706, \textit{p} = 0.04). Furthermore, we find that cine CMR achieves risk discrimination capabilities on par with automated, image-derived phenotypes, removing the dependency on segmentation pipelines. Lastly, we demonstrate that integrating cine CMR with clinical variables through a vision-tabular learning framework stabilizes risk prediction under real-world conditions of incomplete tabular data, a common challenge in clinical practice. Together, these findings position cine CMR as a promising modality for CVD risk assessment.
Bethala, S.; Vanshika,
Show abstract
Automated detection of brain tumors from Magnetic Resonance Imaging (MRI) can accelerate diagnosis and reduce inter-reader variability, yet many existing studies report only top-line accuracy on small datasets, omit efficiency analysis, and provide no interpretability, limiting their clinical credibility. We present a reproducible, comparative, and explainable transfer- learning framework for binary brain-tumor classification. Our framework (i) standardizes a configurable preprocessing pipeline combining CLAHE contrast enhancement and unsharp-mask sharpening, (ii) evaluates a custom CNN baseline and pretrained backbones under an identical training budget, (iii) reports a full metric suite (accuracy, precision, recall, F1, ROC-AUC, PR-AUC, parameter count, and inference latency), and (iv) applies Grad- CAM for spatial interpretability. On a public 253-image MRI dataset (38-image held-out test set), MobileNetV2 achieves the best overall performance (94.74% accuracy, 0.994 ROC-AUC, 0.996 PR-AUC) with only 2.59M parameters and 5.9 ms per- image inference, making it the most deployment-friendly model. Larger backbones (Xception, EfficientNetB0) and the custom CNN converge to degenerate all-positive predictions under the same limited budget, illustrating the small-data overfitting risk that accuracy-only reporting conceals. Grad-CAM confirms that the best model attends to the tumor region. All source code, con- figuration files, and trained evaluation scripts are publicly avail- able at https://github.com/blck-iris/explainable-brain-tumor-mr
Sivakumar, E.
Show abstract
SAM2 (Meta, 2024) provides a strong starting point for segmentation, but given the unique challenges in medical imaging (noise from patient movement, the projection-based nature of X-ray fluoroscopy, and low contrast between vessels and background), direct application is difficult. We fine-tune MedSAM2 on annotated coronary angiograms and apply it to video data for point-of-care use. On the ARCADE validation set (200 images), the fine-tuned model achieves Dice 0.767 compared to 0.033 zero-shot. On 10 fluoroscopic video studies from CoronaryDominance, it tracks vessels coherently and avoids falsely segmenting ribs, stents, and bypass grafts in 9 of 10 studies. Code is available at https://github.com/elakiyasivakumar/SAM2-Coronary-Angiography-VA and the fine-tuned checkpoint at https://huggingface.co/Elakiya17/CA-SAM2.
Tak, D.; Sreedhar, D.; Aerts, H.; Kann, B.
Show abstract
Accurate prediction of tumor recurrence in brain tumor patients following surgery is essential for optimizing adjuvant therapy, response assessment, and surveillance regimen. While MRI remains the gold standard for surveillance, integrating patient-specific clinical context may inform recurrence prediction. Traditional multimodal deep learning approaches often incorporate clinical data via simple fusion, failing to fully capture the semantic interdependencies between visual features and clinical context. Trained on over 5,000 scans from approximately 400 pediatric low-grade glioma subjects and validated across three institutional cohorts, including one clinical trial cohort, our experiments demonstrate incremental performance gains when progressing from vision-only to clinical-vision to a vision-language approach. Our results indicate that converting structured clinical covariates into natural language text allows for more effective synthesis of multimodal data, while providing a platform for incremental addition of clinical context without extending model complexity. We demonstrate that our proposed VLM architecture offers a promising direction for neuro-oncological prognosis by effectively encoding imaging cues and clinical context, with potential applicability to other longitudinal prognosis tasks.
Zhang, C.; Li, H.; Tian, F.; Mansour L., S.; Orban, C.; Chen, C.; Zhou, J. H.; Yeo, B. T. T.; the Alzheimer's Disease Neuroimaging Initiative, ; the Australian Imaging Biomarkers and Lifestyle Study of Ageing,
Show abstract
Longitudinal dementia progression prediction is essential for clinical decision-making. However, models often degrade on external cohorts due to systemic missingness -- where certain biomarkers available during training are completely absent at test time -- compounded by distribution shifts and patient-specific variability. Here, we propose Progression-aware Feature Fusion with Test-Time Adaptation (ProFuse-TTA), a two-stage hierarchical Transformer for longitudinal dementia prediction. Stage 1 learns per-biomarker temporal representations from irregular observations without imputation. Stage 2 fuses them via cross-feature attention, with simulated modality dropout during training for robustness to systemic missingness. At inference, a lightweight test-time adaptation module performs per-individual calibration. We trained on ADNI and evaluated on three external cohorts comprising 2,316 participants and 13,205 timepoints, with controlled modality ablation experiments isolating the effect of systemic missingness. We compared against six baselines, four from a recent benchmark study and two new baselines including one built on a tabular foundation model. ProFuse-TTA achieved the best cross-dataset performance in 8 of 9 settings across clinical diagnosis, MMSE, and hippocampal volume prediction, and ranked first in 14 of 15 ablation scenarios. The model maintained superior performance across varying input lengths and prediction horizons up to 6 years. Pretrained ADNI models are available at XXX.
Jo, A. A.
Show abstract
Maternal healthcare prediction systems often suffer from algorithmic biases due to socio-economic disparities and imbalanced datasets, limiting their effectiveness for equitable healthcare policymaking. This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India. The framework focuses on three critical health indicators:(1) Tetanus Toxoid (TT) booster uptake,(2) immunization coverage rates, and (3) the percentage of pregnant women completing four or more Antenatal Care (ANC) visits. To address fairness, we propose Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training. AESO is model-agnostic and adapts group equity weights in response to real-time disparities. We integrate SHAP, LIME, and feature permutation techniques for explainability, enabling transparent global and local interpretation. Empirical results demonstrate that MaternaAI significantly improves fairness metrics and model accuracy across diverse machine learning and deep learning models, offering interpretable and equitable decision support for public health stakeholders.
Levitis, E.; Tregidgo, H. F. J.; Zimmerman, D.; Jung, B.; Karandikar, S.; Gardner, M.; Mattisson, P.; Kafadar, E.; Zapaishchykova, A.; Kann, B. H.; Sotardi, S. T.; Vossough, A.; Huang, H.; Billot, B.; Iglesias Gonzales, J. E.; Alexander, D. C.; Alexander-Bloch, A. F.; Seidlitz, J.
Show abstract
Clinical brain MRIs from pediatric health systems represent a viable resource for modeling early neurodevelopmental trajectories and studying neurodevelopmental risk in real-world populations. However, a limitation to date has been the performance of existing segmentation tools for measuring various brain phenotypes in clinical scans. In particular, many tools underperform in infant scans due to morphological and physical changes such as rapid myelination. Here, we introduce ClinSeg: a robust segmentation approach tailored to early-life clinical MRIs with variable orientation, resolution, and contrast. We leverage existing registration and synthetic data generation tools to construct a training corpus for a 3d U-Net spanning anatomical and contrast diversity, including scans with morphological abnormalities from a pediatric hospital. Validated against manual segmentations, ClinSeg outperforms existing models in infancy while matching them in childhood and adolescence. Finally, ClinSeg enables the construction of reference brain growth trajectories in 11,699 individuals from 0-21 years of age, leading to the detection of more nuanced age-related findings in clinical groups.
Di Giovanni, D. A.; Takada, A.; McNabb, E.; Dana, J.; Yokota, H.; Tsuboyama, T.; Zakarian, R.; Vallieres, M.; Tsui, J. M. G.; Reinhold, C.
Show abstract
Purpose: To evaluate how segmentation architecture and dataset-adaptive configuration influence uterine MRI segmentation across heterogeneous benign and malignant tasks. Methods: U-Net, Swin-UNETR, and MedNeXt were compared with nnU-Net as a self-configuring reference across T2-weighted MRI datasets: public multiclass UMD anatomy/fibroid segmentation (n=300), institutional endometrial cancer tumor segmentation (n=206), and institutional uterine mass lesion segmentation (n=234). A relabeled external UMD-style cohort (n=12) assessed domain shift. Models used fixed partitions, fold ensembling, Dice, HD95, ASSD, volume error, and paired bootstrap comparisons with Holm correction. Results: MedNeXt was the strongest manually controlled architecture. nnU-Net achieved the highest performance on all internal datasets and external testing. Macro-Dice reached 0.761, 0.746, and 0.814 for nnU-Net on UMD, endometrial cancer, and uterine mass datasets, respectively, versus 0.722, 0.726, and 0.789 for MedNeXt. The nnU-Net-MedNeXt gap was largest for multiclass UMD segmentation and smaller in binary tasks. External testing degraded all models; nnU-Net remained highest (0.542), followed by MedNeXt (0.490), U-Net (0.396), and Swin-UNETR (0.287). Conclusions: Uterine MRI segmentation performance depended on task, architecture, and evaluation domain. MedNeXt supported modern convolutional design as a strong manual baseline, but nnU-Net remained the most robust overall, emphasizing the importance of dataset-adaptive configuration and external validation.
Amiri, S.; Afshar, P.; Rohban, M. H.
Show abstract
Objectives. Radiomics pipelines extract hundreds of quantitative features that are widely known to be redundant, but the structure of this redundancy is usually treated as a per-dataset nuisance to be pruned away. We tested the alternative hypothesis that a substantial number of feature-feature correlations are universal: they persist across patients and across anatomically distinct structures because they reflect shared mathematical and image-statistical properties of how the image is summarised, rather than properties of the tissue being imaged. Materials and Methods. We re-analysed the publicly available Radiomics Atlas Dataset of normal Abdominal and Pelvic CT (RADAPT), restricting the analysis to the 526 non-contrast-enhanced examinations of the 531-subject atlas and to the 107 original (non-filtered) PyRadiomics features. The 53 segmented structures were grouped into four broad anatomical categories -- bones, muscles, vessels, and parenchymal organs. RADAPT is distributed as one Excel file per structure, with patients as rows and features as columns. Within each structure file we z-score-normalised every feature across patients, computed the absolute Spearman correlation matrix, and retained edges with |{rho}| [≥] {tau} for {tau} in {0.70, 0.80, 0.90}. We then intersected the edge sets across all structure files to obtain a "universal" correlation graph, in which an edge survives only if it exceeds the threshold in every structure (each estimated across the full patient sample). Stable feature communities were defined as the maximal cliques of this graph. Robustness to patient sampling was tested by repeating the entire pipeline on five independent random splits of each file into two patient halves (10 sub-cohorts per threshold), and the implementation was independently reproduced in R. Results. Despite the strictness of the global-intersection criterion, 34, 24, and 14 stable feature communities survived at {tau} = 0.70, 0.80, and 0.90 respectively, with the largest cliques containing six members at {tau} = 0.70 and {tau} = 0.80 and five members at {tau} = 0.90. The community structure was clearly interpretable: separate cliques captured (i) variance-like intensity dispersion, (ii) long-run / low-frequency (coarse) texture, (iii) high gray-level texture, (iv) low gray-level texture, (v) volume and surface shape, and (vi) local-homogeneity and energy/entropy duals. On random-half resampling the exact-match recovery rate of these communities was 81.5 %, 86.7 %, and 80.7 % across the three thresholds; departures from exact recovery were almost always a single boundary feature added or dropped, consistent with finite-sample fluctuation of near-threshold edges rather than structural instability. The R re-implementation reproduced the Python results exactly. Conclusion. A substantial portion of radiomics feature collinearity is universal across patients and tissues. We distinguish two layers within it: trivial near-algebraic duals that are universal by construction, and non-trivial cross-matrix-family communities that are the genuine empirical finding. Together they provide an interpretable, definition-grounded basis for aggressive dimensionality reduction, for retrospectively reconciling apparently different feature selections in the literature, and for moving radiomics pipelines toward organ-agnostic, more reproducible models. Clinical relevance statement. Selecting a single representative feature from each universal community shrinks the original-feature space by roughly an order of magnitude without sacrificing biologically distinct information. For example, the five variance-family members (first-order Variance, GLCM SumSquares, GLCM ClusterTendency, GLDM and GLRLM GrayLevelVariance) can be replaced by a single representative, removing redundant degrees of freedom that would otherwise inflate model variance; and labelling each retained feature by its community lets two studies that selected different variance-family names be recognised as having found the same signal, simplifying model development and improving cross-cohort generalisability in clinical CT workflows.
Salis, F.; Tallone, N.; Fina, P. R.; Massobrio, R.; Bellacosa Marotti, R.; Conti, D.; Fuso, L.; Mariani, L.; Ferrero, A. M.; Accomasso, F.; Arena, A.; Borella, F.; Casula, V.; Cosma, S.; De Grandis, T.; Grisaru, D.; Lacalandra, A.; Pereira Sanchez, A.; Seracchioli, R.; Robba, E.; Roccio, M.; Gerace, F.
Show abstract
Ovarian cancer is recognized as the deadliest gynecological malignancy. Diagnosis at advanced stages and the lack of effective screening program lead to poor survival rates, dropping to 17-39 % in stage III-IV diseases. Ultrasound (US) is the primary imaging modality for ovarian structures evaluation, but it is strongly affected by the operator expertise due to the complexity of adnexal masses and the physiological variability of ovarian morphology throughout a womans lifecycle. The International Ovarian Tumor Analysis (IOTA) group introduced definitions and predictive tools to standardize gynecological US interpretation. However, these tools still rely on subjective interpretation, thus highlighting the need for more objective solutions. Recent studies have explored artificial intelligence (AI) algorithms for gynecological US, mainly focusing on adnexal masses classification. Conversely, a robust solution supporting the identification and description of healthy and tumoral ovarian structures is still lacking. This paper proposes OvAi Focus, a framework including (i) a segmentation module for the identification of healthy ovaries, plus solid and cystic components of adnexal masses; (ii) a morphology module for the extraction of IOTA-based keywords to describe adnexal masses morphology. Segmentation module results were compared to ground truth masks, showing DICE scores from 0.62 for functional ovary to 0.87 for the whole adnexal mass. Morphology module was tested through interobserver agreement analysis, obtaining Fleiss Kappa from 0.16 to 0.57 and Percent Agreement from 47 to 90 %, in line with existing literature. OvAi Focus represents an innovative solution which could help overcoming subjectivity in gynecological US imaging interpretation.